Papers with safety rate

2 papers
Athena: Safe Autonomous Agents with Verbal Contrastive Learning (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing safety benchmarks on the ability of large language models to perform tasks are lacking.
Approach: They propose a framework that leverages verbal contrastive learning to guide agents towards safety . they use past safe and unsafe trajectories as in-context examples to guide them towards safety.
Outcome: The proposed framework leverages verbal contrastive learning to guide agents towards safety while performing tasks.
PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient .
Approach: They propose a framework that explicitly models personalized risk inference and memory evolution.
Outcome: The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations